Skip to main content

AI Architecture

This page is about where AI fits in a health architecture — the components, the data flows, the controls. The clinical, regulatory and evaluation questions are covered in clinical AI, machine learning and AI ethics, and they are not optional reading: an AI system deployed without them is a patient safety risk regardless of how well it is engineered.

A note on maturity. Standards in this area are thin. FHIR, CDS Hooks and SMART are Tier 1 and stable. Most AI-specific integration patterns, including MCP, are Tier 4 — emerging, changing, and not health interoperability standards. This page marks which is which.


Where AI attaches to a health system​

┌──────────────────────────────────────────────────────────┐
│ Point of care │
│ EMR ──CDS Hooks──▶ decision service ──▶ card in UI │ ← real-time,
│ EMR ──SMART app──▶ model-backed application │ in workflow
└──────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────┐
│ Asynchronous │
│ Event ──▶ inference service ──▶ Observation / Flag │ ← risk scores,
│ Imaging ──▶ AI ──▶ DICOM SR / segmentation │ triage
└──────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────┐
│ Population / batch │
│ Bulk export ──▶ pipeline ──▶ cohort, risk strata │ ← planning,
│ │ surveillance
└──────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────┐
│ Assistive / knowledge │
│ Clinician question ──▶ RAG over guidelines ──▶ answer │ ← reference,
│ with citations │ not diagnosis
└──────────────────────────────────────────────────────────┘

The four attachment points have very different risk profiles. Real-time decision support at the point of care is the highest risk and the most regulated; retrieval over published guidelines with citations is the lowest, and is where most organisations should start.


Reference components​

An AI platform in a health ecosystem needs more than a model server.

ComponentPurpose
Feature / data pipelineAssembles inputs from the shared health record or warehouse, reproducibly
Model registryVersioned models with lineage: training data, code, parameters, evaluation results, approver
Inference serviceServes predictions with an API contract, versioning and latency guarantees
MonitoringInput drift, output distribution, performance against outcomes, subgroup breakdowns
Human-in-the-loop interfaceHow a clinician sees, questions and overrides the output
AuditEvery inference recorded: model version, inputs, output, who saw it, what they did
Governance registryWhich models are approved for which use, by whom, with what conditions and review date
Feedback loopOutcomes flowing back for evaluation — the part that is almost always missing

The audit requirement deserves emphasis. When a clinical decision is questioned months later, you must be able to reconstruct exactly what the model said, which version said it, on what inputs, and whether the clinician followed it. Design this in; it cannot be added retrospectively.


Integration patterns​

CDS Hooks + inference service (Tier 1 integration, real-time)​

The EMR calls out at a defined workflow point; the service returns cards. Standards-based, vendor-neutral, and the recommended path for real-time decision support. See SMART on FHIR.

The constraint is latency: the clinician is waiting. Budget a few hundred milliseconds, which rules out large-model inference in the synchronous path unless it is cached or precomputed.

SMART app with model backend (Tier 1 integration)​

A full application launched from the EMR with patient context. Suits richer interaction — explanation, exploration, structured input — where a card is too small.

Event-triggered scoring (Tier 1 integration, asynchronous)​

A subscription or event triggers scoring; the result is written back as a FHIR Observation, RiskAssessment or Flag. No latency constraint, and the output lands in the record where it can be found later.

Write model outputs to the record clearly labelled as model-derived, with the model version in Provenance. A risk score indistinguishable from a clinician's assessment is a safety problem and a data-quality problem — it will end up as a feature in the next model.

Batch over bulk export (Tier 1 integration)​

Bulk Data to a pipeline, for population risk stratification, quality measurement or research. Requires the same governance as any bulk disclosure.

RAG over clinical knowledge (Tier 4)​

Retrieval-augmented generation over a curated corpus — national guidelines, formularies, protocols — with citations.

Question
│
▼
Retrieval ──▶ curated corpus (guidelines, formulary, protocols)
│ versioned, with effective dates
▼
Generation ──▶ answer + citations to specific source passages
│
▼
Clinician verifies against the cited source

Design points that matter more in health than elsewhere:

  • Corpus governance is the whole system. RAG over an unmaintained document pile produces confident answers from superseded guidance. Every document needs an owner, a version and an effective date, and superseded documents must be removed or clearly marked.
  • Citations must be specific and checkable — the passage, not the document. The clinician verifying the citation is the safety control.
  • Terminology-aware retrieval improves recall substantially: expanding a query through SNOMED CT synonyms and hierarchy finds documents that lexical search misses.
  • Scope it to reference, not to patient-specific advice. "What does the national protocol say about magnesium sulphate in eclampsia?" is a retrieval question. "Should this patient receive it?" is a clinical decision, and moving from one to the other crosses a regulatory line — see clinical AI.

Agentic workflows (Tier 4, and the most cautious ground here)​

Systems that plan and take actions across tools. In a health context the distinction that matters is between reading and acting:

  • Read-only agents that gather and summarise information for a human are a manageable risk with good audit
  • Agents that write to clinical records, place orders or send messages to patients require the same scrutiny as any automated clinical action, and should not be deployed without regulatory analysis, human confirmation of every consequential action, and a full audit trail

The technology is moving quickly and the governance is not. Be explicit about which category any deployment falls into.


Data for AI​

  • Provenance and consent. What is the legal basis for using this data to train? Consent for treatment is generally not consent for model development. See consent and trust.
  • De-identification is a deliberate step with residual re-identification risk, not a checkbox. See health data.
  • Representativeness. A model trained on tertiary hospital data will underperform in primary care and in rural settings — and the underperformance will fall on the populations already least well served. Report performance by subgroup, or you have not evaluated the model.
  • Terminology normalisation. Data coded inconsistently across facilities produces features that encode the facility rather than the patient.
  • Label quality. Health "labels" are usually recorded outcomes, which are themselves shaped by who had access to care.

Monitoring in production​

A model that performed well at validation will degrade. Monitor:

  • Input drift — the population, coding practice or referral pattern has changed
  • Output drift — the distribution of predictions has moved
  • Performance against outcomes, where outcomes are observable, with a defined lag
  • Subgroup performance — by sex, age, geography, facility type, and any locally relevant equity dimension
  • Clinician override rate — a rising override rate is the earliest available signal that something is wrong, and it is free to collect
  • Downstream effect — did the intervention the model triggers actually change anything?

Define in advance what triggers withdrawal of the model, and who has authority to withdraw it. A model with no defined stopping rule stays in production.


Governance​

Minimum before any model influences care:

  • Named clinical owner, accountable for its use
  • Documented intended use, and explicit out-of-scope uses
  • Evaluation on local data, reported by subgroup
  • Regulatory position established — is it a medical device in this jurisdiction? See clinical AI
  • Legal basis for the data used in development
  • Human-in-the-loop design, with override always available and recorded
  • Monitoring plan with thresholds and a stopping rule
  • Audit of every inference
  • Scheduled review date
  • Incident process for suspected model harm

See AI ethics and governance.


In this section​


References​